Resumen:
Conventionally, synthetic training data quality is evaluated through human perception, prioritizing visual realism. From the model’s perspective, what truly matters is whether a sample lies within the right region of its embedding space. This work introduces VERSE, a methodology for analyzing and improving the performance of Vision–Language Models by exploring their visual embedding space. VERSE enables the visualization of latent representations to assess model feasibility, identifies problematic regions, and guides synthetic data generation to enhance performance in those clusters. We validate the proposed methodology for Visually-rich Document Understanding by training on the synthetic MERIT Dataset and evaluating on its real-world counterpart, MERIT Secret, focusing on key information extraction as a sequence-generation task scoped to transcripts of records in Spanish. Results show that VERSE uncovers the visual features associated with error-prone clusters, and that retraining with samples containing these features substantially boosts F1 performance without degrading generalization. On-premise models optimized with VERSE—Donut (F1 = 0.76) and Idefics2 (F1 = 0.81)—match or surpass SaaS solutions such as GPT-4o (F1 = 0.78) and Pixtral (F1 = 0.73), preserving data privacy and avoiding external APIs.
Resumen divulgativo:
VERSE es una metodología para analizar y mejorar VLMs explorando su espacio de embeddings visuales. Identifica clústeres propensos a error y guía la generación de datos sintéticos, logrando que modelos on-premise (Donut, Idefics2) igualen o superen a soluciones SaaS como GPT-4o preservando la privacidad.
Palabras Clave: Visually-rich Document Understanding; Vision-Language Models; Visual embeddings; Interpretability; Explainability
Índice de impacto JCR-JIF y cuartil WoS: 9,100 - Q1 (2025)
Referencia DOI:
https://doi.org/10.1016/j.patcog.2026.114448
Publicado en papel: Diciembre 2026.
Publicado on-line: Julio 2026.
Cita:
I. de Rodrigo, A.J. López López, J. Boal, "VERSE: Visual Embedding Reduction and Space Exploration - Latent-space clustering for improving document understanding", Pattern Recognition, Vol. 180, nº. Part D, pp. 114448, Diciembre 2026. [Online: Julio 2026] doi: 10.1016/j.patcog.2026.114448